European Heart Journal - Digital Health
◐ Oxford University Press (OUP)
Preprints posted in the last 90 days, ranked by how well they match European Heart Journal - Digital Health's content profile, based on 18 papers previously published here. The average preprint has a 0.04% match score for this journal, so anything above that is already an above-average fit.
Aminorroaya, A.; Vasisht Shankar, S.; Carter, M.; Khan, M.; Dhingra, L. S.; Khunte, A.; Croon, P. M.; Lombo, B.; McNamara, R. L.; Oikonomou, E. K.; Pedroso, A. F.; Khera, R.
Show abstract
Importance: Consumer wearables such as the Apple Watch can record single-lead electrocardiograms (ECGs) but are used mainly to detect rhythm disorders. Artificial intelligence-enhanced ECG (AI-ECG) could extend these real-world recordings for detecting structural heart disease (SHD), yet prospective validation remains limited. Objective: To prospectively validate a previously developed, noise-adapted AI-ECG model for detecting severe SHD from single-lead Apple Watch ECGs. Design: Prospective cohort study. Setting: Yale New Haven Hospital echocardiography laboratory. Participants: Adults aged >=18 years undergoing outpatient transthoracic echocardiography (TTE) as part of routine clinical care. Exposure: A 30-second, single-lead Apple Watch ECG recorded during the TTE visit and processed through an end-to-end, HIPAA-compliant platform for real-time AI-ECG inference. Main Outcomes and Measures: The primary outcome was discrimination for TTE-defined severe SHD, a composite of left ventricular systolic dysfunction (left ventricular ejection fraction <40%), severe left-sided valvular disease, and/or severe left ventricular hypertrophy, assessed by the area under the receiver operating characteristic curve (AUROC). Secondary measures were sensitivity, specificity, negative predictive value (NPV), and positive predictive value (PPV) at prespecified thresholds, and screening efficiency, assessed by the number needed to test (NNT) under usual-care versus AI-ECG-guided strategies. Results: Among 596 participants with analyzable Apple Watch ECGs (median age, 62 years [IQR, 46-72]; 51.2% women), severe SHD was present in 30 (5.1%). The model discriminated severe SHD well (AUROC, 0.841; 95% CI, 0.761-0.921), with a sensitivity of 76.7% (59.1-88.2), specificity of 83.2% (79.9-86.1), NPV of 98.5% (97.0-99.3), and PPV of 19.7% (13.5-27.8) at the prespecified threshold. An AI-ECG-guided strategy reduced the NNT to identify one case by more than 60% versus usual care across the composite and individual SHD phenotypes. Conclusions and Relevance: In this prospective cohort, a noise-adapted AI-ECG algorithm identified SHD phenotypes from real-world single-lead Apple Watch ECGs and improved screening efficiency. These findings support a potential role for wearable ECG-based screening in the scalable identification of clinically actionable SHD.
Meneguitti Dias, F.; Ribeiro, E.; Olivetti, N.; Carvalho, O.; Krieger, J. E.; Gutierrez, M.
Show abstract
Automated electrocardiogram analysis has advanced largely through digital waveforms, yet many emergency-care workflows rely on ECGs available only as printed tracings, scanned reports, PDFs or mobile photographs. We developed an image-based deep learning system for emergency ECG classification and evaluated it in InCor-EMG, an expert-adjudicated dataset of 18,519 emergency ECGs spanning 12 ECG categories, with labels from 19 cardiologists. On the held-out test set, the final ConvNeXt ensemble achieved a macro F1-score of 0.807 (95% CI, 0.788-0.825), compared with 0.820 (95% CI, 0.805-0.832) for annotating cardiologists, and higher F1-scores than Mortara Veritas in most evaluated categories. Performance was associated more strongly with inter-reader agreement than with training sample size and remained informative across scanned and photographed ECGs, with supportive performance in model-enriched temporal and heterogeneous public-image evaluations. These findings support ECG image classification when digital waveforms are unavailable.
Pitre, T.; Marques, L.; Weatherald, J.; Mak, S.; Thavendiranathan, P.; Granton, J.
Show abstract
Background: Right ventricular (RV) function predicts survival in pulmonary hypertension (PH) and other cardiovascular diseases, yet echocardiographic AI has largely focused on the left ventricle (LV). Objectives: To develop and evaluate PH-ECHO-AI, a unified deep learning model performing four-chamber segmentation, landmark localisation, biventricular ejection fraction (EF) estimation, deformation analysis, and PH prediction from a single apical four-chamber (A4C) clip. Methods: We developed the model using 8,416 clips from four public datasets and no institutional data: EchoNet-Dynamic, CAMUS, RVENet (apical four-chamber clips paired with 3D-echocardiographic right ventricular ejection fraction, RVEF), and MIMIC-IV-ECHO. Evaluation used held-out, training-excluded data with expert-reviewed reference standards and a per-cohort audit of patient-level separation: 1,416 clips for segmentation; 600 clips for function and deformation (350 referenced to 3D-echocardiographic RVEF, 250 to the EchoNet LVEF); and 1,076 MIMIC-IV patients for PH prediction, with five-fold cross-validation. Performance measures were Dice, correlation, mean absolute error (MAE), Bland-Altman agreement, and area under the receiver operating characteristic curve (AUC). Results: Four-chamber segmentation generalised robustly across all datasets (pooled Dice: LV 0.925, RV 0.836, LA 0.910, RA 0.904). Left ventricular ejection fraction (LVEF) was estimated with r=0.845 (95% CI 0.786 to 0.886) and MAE 4.67%. RVEF, regressed directly from the clip by a supervised head trained on 3D-echocardiographic labels with no geometric assumption, reached r=0.754 (95% CI 0.690 to 0.806) and MAE 4.98%, matching published single-view RVEF ceilings and exceeding geometric RV fractional area change (RVFAC; r=0.278). Deformation and excursion metrics, namely RV free-wall and LV A4C longitudinal strain and tricuspid and mitral annular plane systolic excursion (TAPSE, MAPSE), proved physiologically coherent. Segmentation generalised to the external MIMIC-IV cohort, and PH prediction was developed and evaluated entirely within it; RVEF evaluation was clip-disjoint and same-source, so cross-centre RVEF validation remains outstanding. Using echocardiographic geometry alone, confirmed PH was detected with an AUC of 0.697 and strong calibration (Brier 0.061). Conclusions: A single, reproducible model provides comprehensive right-heart-focused interpretation from one A4C view. It achieves RVEF accuracy competitive with dedicated RV models while simultaneously delivering segmentation, deformation, annular excursion (TAPSE and MAPSE), and PH prediction. Registration: This retrospective study used existing datasets. Code is openly released, and trained model weights are available to credentialed investigators, for independent evaluation.
Nicolson, A.; Pröll, S.; Lunelli, R.; Blankenburg, H.; Pramstaller, P.; Fuchsberger, C.; Bauer, A.; Dlaska, C.
Show abstract
Background Recent artificial intelligence (AI) models applied to the electrocardiogram (ECG) for risk stratification typically rely on supervised learning, defining risk as the error relative to an external target such as age or sex. This couples the risk score to the choice of target rather than the cardiac signal alone, and may limit generalisability. We aimed to develop a self-supervised AI-ECG risk score based on the error in reconstructing a partially masked ECG. Methods A transformer-based masked autoencoder was trained on 85% of the CODE dataset (n = 7,212,109 ECGs) to reconstruct ECG signals from partially masked inputs. The association between reconstruction error and all-cause mortality was assessed internally in CODE-15% and externally validated in four independent cohorts: MIMIC-IV-ECG (critical care, US), HEEDB (hospital, US), CHRIS (population-based, Italy), and Innsbruck (cardiology centre, Austria). A binary risk score (>1 SD above the CODE-15% mean) was additionally evaluated in these cohorts and in the UK Biobank (population-based, UK). Findings In Cox proportional hazards models adjusted for age and sex, each 1-SD increase in reconstruction error was associated with higher all-cause mortality (all p<0.001; cohort median follow-up 1.4-11.0 years): CODE-15% (HR 1.39, 95% CI 1.37-1.42), MIMIC-IV-ECG (HR 1.39, 95% CI 1.37-1.40), HEEDB (HR 1.41, 95% CI 1.40-1.41), Innsbruck (HR 1.23, 95% CI 1.21-1.26), and CHRIS (HR 1.25, 95% CI 1.14-1.38). The binary threshold identified a high-risk group with increased mortality in all six cohorts, including the UK Biobank (HR 1.27, 95% CI 1.08-1.50, p=0.004). Interpretation Reconstruction error is a generalisable predictor of all-cause mortality across diverse clinical and population-based settings. Unlike supervised approaches, it reflects the model's uncertainty about the ECG signal itself rather than error relative to an external target, providing a direct measure of how much each recording deviates from normal cardiac electrical patterns.
Aydogdu, D.; Gaber, F.; Sorooshmehr, A.; Akalin, A.
Show abstract
Cardiovascular diseases (CVDs) remain the primary global health burden, motivating the search for robust, non-invasive risk biomarkers. We harness a foundation model pretrained on over 10 million recordings, to evaluate ECG-derived age deviation as a cross-cohort biomarker of CVD burden. A predictive model, trained exclusively on healthy subjects, achieved accurate age prediction. Diseased subjects exhibited significant positive age acceleration across multiple categories, with structural and ischemic heart diseases showing the largest effects. External validation in a hospital-based cohort (n=160,493) confirmed that age acceleration independently predicts all-cause mortality, with the strongest prognostic value in patients under 65 years. Furthermore, we demonstrated that disease discrimination and mortality prediction are preserved across 6-lead and single-lead configurations, supporting potential deployment in wearable or mobile devices. Our analysis also revealed a striking morphological confound from the complete left bundle branch block, leading us to propose absolute age deviation as a more robust, universal risk marker. These findings establish ECG-derived biological age deviation as a highly generalizable and clinically actionable biomarker for assessing cardiovascular risk. We have also developed a web application at https://bioinformatics.mdc-berlin.de/ECGage that allows users to easily test our framework.
Sriram, R.; Nenadic, I.; Shahrabani, E.; Goonewardena, S.; Yao, S.; Farrell, B.; Loring, Z.; Murthy, V. L.
Show abstract
We conducted a scaling evaluation of unlabeled pretraining for electrocardiogram foundation model performance. One-dimensional vision transformer masked autoencoders were pretrained across increasing ECG volumes and fine-tuned for rhythm, morphology, diagnostic, and structural heart disease tasks. Models pretrained below 400,000 ECGs failed to consistently exceed controls without self-supervised pre-training, whereas 600,000 to 800,000 ECGs improved AUROC across tasks, suggesting a minimum threshold for effective ECG representation learning.
Pan, L.; Li, S.; Huo, J.; Xiao, Z.; Yu, Z.; Chen, J.; Zhou, Y.; Li, Z.; Zhang, B.; Li, X.; Wang, C.; Lu, H.; Patlatzoglou, K.; Kramer, D. B.; Waks, J. W.; Ng, F. S.; Liang, Y.; Ge, J.
Show abstract
Background: Heart failure with reduced ejection fraction (HFrEF) remains a major global health burden. Most electrocardiogram (ECG)-based artificial intelligence models are limited to diagnostic tasks or fixed-horizon prognostic classification and provide little insight into the temporal evolution of risk. In addition, concerns regarding model interpretability continue to impede clinical adoption. Whether deep learning applied to ECGs can deliver individualized, time-resolved, and biologically interpretable risk estimates for incident HFrEF across diverse populations remains uncertain. Methods: We developed a convolutional neural network-based survival model using raw 12-lead ECGs from Zhongshan Hospital (SHZS) and externally validated it in independent cohorts from Shanghai Tenth People's Hospital (SHTP) and Beth Israel Deaconess Medical Center (BIDMC). The model generated individualized, day-by-day probabilities of incident HFrEF over a 5-year horizon. Performance was comprehensively evaluated using discrimination, calibration, precision-recall characteristics, clinical utility, and risk stratification metrics, with subgroup analyses across age, sex, and race to assess generalizability. Model interpretability was examined using complementary representation and attention-based frameworks. Results: In 458,884 patients, the survival model demonstrated strong and stable discrimination across cohorts, with overall C-indices of 0.971 (95% CI, 0.965-0.976) in SHZS, 0.945 (95% CI, 0.938-0.950) in SHTP, and 0.855 (95% CI, 0.850-0.860) in BIDMC, and consistently high time-dependent AUROC values across the 1-5-year horizons. Calibration showed close agreement between predicted and observed risks, and decision curve analyses indicated meaningful net clinical benefit across a broad range of thresholds. Kaplan-Meier curves showed clear stratification across predicted risk groups. Interpretability analyses identified physiologically coherent ECG features related to QRS duration, heart rate, and QT interval that were associated with predicted risk. Conclusion: This ECG-based deep learning survival model provides individualized, time-resolved, and clinically interpretable estimates of future HFrEF risk with robust performance across multinational cohorts. These findings support the potential of AI-enabled ECG analysis as an accessible tool for early HFrEF risk stratification within routine clinical workflows.
Li, S.; Zhang, B.; Pan, L.; Xiao, Z.; Yu, Z.; Chen, J.; Zhou, Y.; Li, Z.; Li, X.; Wang, C.; Lu, H.; Lai, H.; Liang, Y.; Ge, J.
Show abstract
Background: Regurgitant valvular heart disease (rVHD) is a major cause of cardiovascular morbidity. Echocardiography is the diagnostic standard but is resource-intensive for large-scale screening. Electrocardiography (ECG) has shown promise for predicting incident rVHD, yet performance varies across phenotypes, particularly for aortic regurgitation (AR). Chest radiography (CXR) provides complementary structural and hemodynamic information. We hypothesized that a multimodal model integrating ECG and CXR would improve prediction of incident moderate-to-severe rVHD. Methods: In this retrospective multicenter study, we identified 212,888 paired ECG-CXR examinations from 116,380 patients across two Chinese centers. Baseline ECG and CXR were obtained within 60 days of echocardiography. Outcome was progression to moderate-to-severe AR, mitral regurgitation (MR), or tricuspid regurgitation (TR). We developed a multimodal neural network with pretrained unimodal encoders, token-level cross-modal fusion, and a class-specific gating mechanism that adaptively weighted ECG-only, CXR-only, and fused predictions. Performance was assessed using C-index, AUROC, AUPRC, decision curve analysis, net reclassification improvement (NRI), and Kaplan-Meier stratification. Results: Multimodal fusion consistently outperformed unimodal models across all phenotypes. For AR, C-index improved from 0.616 (ECG-only) to 0.713 (multimodal; AUROC 0.729, AUPRC 0.972). For MR, multimodal C-index was 0.801 (AUROC 0.814, AUPRC 0.972), versus 0.782 for ECG and 0.775 for CXR alone. For TR, multimodal and CXR-only models showed similar discrimination (C-index 0.802), but multimodal fusion yielded greater net benefit on decision curve analysis. NRI was positive across all time horizons (1-5 years) for all valve types. Grad-CAM interpretability analyses revealed that ECG attention localized to leads II, V-V (AR), leads I, II, aVF, V-V (MR), and inferior/right precordial leads (TR); CXR attention highlighted chamber-specific enlargement and pulmonary congestion patterns consistent with pathophysiology. Conclusion: A multimodal deep learning model integrating ECG and CXR significantly improved prediction of incident rVHD compared with ECG alone, with the greatest benefit observed for AR. The model leveraged complementary electrical and structural information, demonstrated biological plausibility through interpretability analyses, and provided consistent clinical utility. Given the widespread availability and low cost of both modalities, this approach offers a scalable tool for risk stratification in routine care. Prospective studies are warranted to validate clinical implementation.
Ekambarapu, L.; Pendyal, A.; Lin, A.; Alwakeel, M.; Rajaratnam, A.
Show abstract
Background: Unstructured biomedical data, such as echocardiography reports, are rich in information but time consuming to analyze at scale. Rule-based, regular expression-driven terminology mapping can only extract individual variables while large language models (LLMs) offer scalable and clinically meaningful interpretations of heterogeneous disease processes. Right ventricular dysfunction (RVD) is an example of a multifactorial disease state in which key structural and physiologic features are captured both narratively and in structured fields, making it an ideal test case for evaluating whether LLMs can recover complex phenotypes that rules based methods routinely miss. Purpose: To compare an LLM-based extraction method to a conventional rules-based schema for identifying and phenotyping echocardiographic features associated with RVD in a large TTE dataset. Methods: MIMIC-III NOTE2NUM echocardiography reports (n = 45,794) were analyzed using GPT-4o-based LLM extraction deployed within a secure health system enclave and were benchmarked against echocardiographic measurements defined in the MIMIC-III dictionary schema. In MIMIC-III, PH was recorded qualitatively (mild/moderate/severe) based on tricuspid regurgitant (TR) jet velocity and then re-coded as present vs. absent. LLM based extraction defined RVD as (1) RV structural abnormality (>= 1 of hypertrophy, dilation, or wall hypo-/akinesis) or (2) RV pressure/volume overload (>= 2 of the following: estimated right atrial pressure > 8 mmHg, TR jet velocity > 2.8 m/s, fractional area change < 35%, tricuspid annular planar systolic excursion < 17 mm, S' < 9.5 cm/s, or E/e' > 14), with PH defined as estimated pulmonary artery systolic pressure > 35 mmHg or qualitative documentation of PH. Results: LLM extraction identified PH in 15,394 (33.6%), RV pressure/volume overload in 14,449 (31.6%), and RV structural abnormalities in 11,955 (26.1%). Co-occurrence was common: overload + structural changes in 9,380 (20.5%), overload + PH in 9,756 (21.3%), structural changes + PH in 6,183 (13.5%), and all three in 5,620 (12.3%). Using the MIMIC-III dictionary schema, PH prevalence was similar (15,371; 33.6%), but RV overload fields were captured less often (pressure overload 1,357 [3.0%], volume overload 1,128 [2.5%], pressure + volume overload 1,093 [2.4%]; any overload field 3,578 [7.8%]), and RV pressure/volume overload with PH was identified in only 731 (1.6%). Conclusions: LLM-based extraction outperforms rules-based schemas for identifying complex disease states not defined by any single variable. By synthesizing multifactorial signals, LLMs can phenotype RVD with higher fidelity and support population-level assessment. Further validation using multimodality imaging, invasive hemodynamics, and clinical outcome data is needed.
Li, Z.; Sun, Y.; Jiang, C.; Pan, T.; Zhou, Y.; Wang, C.; Pan, L.; Zhang, X.; Yang, Z.; Yu, Z.; Xiao, Z.; Chen, J.; Huang, Y.; Sun, R.; Gan, Y.; Li, X.; Zhang, B.; Zhang, Z.; Wang, X.; Han, L.; Qi, Y.; Cheng, Y.; Liang, Y.; Ge, J.
Show abstract
BACKGROUND: Coronary angiography remains the reference standard for diagnosing coronary artery disease and guiding revascularization, yet its interpretation requires expert integration of multi-view anatomy, lesion morphology and procedural context. Existing artificial intelligence approaches are largely task-specific, annotation-dependent and limited in capturing the semantic relationship between angiographic findings and interventional decision-making. Whether large-scale vision-language pretraining can enable transferable foundation-model representations for invasive coronary imaging remains unknown. METHODS We developed CAG-MIND, a domain-specific vision-language foundation model for coronary angiography, using 135,475 CAG examinations paired with procedural reports, comprising 812,850 angiographic videos from Zhongshan Hospital and Shanghai Geriatric Medical Center. Each case consisted of standardized six-view angiographic acquisitions paired with structured procedural semantics extracted from routine reports using a large language model-assisted pipeline. The model was pretrained by aligning multi-view angiographic representations with report-derived semantic embeddings through bidirectional contrastive learning. Performance was evaluated under zero-shot and supervised fine-tuning settings across 11 downstream tasks grouped into structural abnormality detection, atherosclerotic plaque assessment, and interventional decision prediction, using both an internal validation cohort and an independent external test cohort. RESULTS CAG-MIND demonstrated robust performance across all three task categories. In the zero-shot setting, the model achieved mean AUROCs of 0.686 in the internal validation cohort and 0.745 in the external test cohort, indicating transferable multimodal representations without task-specific supervision. Following supervised fine-tuning, the mean AUROC increased to 0.827 and 0.846, respectively, with excellent performance for coronary stenosis detection (AUROC 0.940 in both cohorts), balloon/stent prediction (0.900 and 0.907), and CABG recommendation (0.877 and 0.875). Compared with representative biomedical vision-language models and conventional image-based architectures, CAG-MIND consistently achieved superior performance in both zero-shot and supervised settings and remained superior to fully fine-tuned competing models when trained with only 10% of the labelled data. Grad-CAM visualization demonstrated anatomically plausible lesion-focused attention, supporting the interpretability of the learned representations. CONCLUSIONS CAG-MIND is, to our knowledge, the first large-scale vision-language foundation model for coronary angiography trained at more than 100,000-patient scale. By aligning standardized multi-view angiographic videos with report-derived procedural semantics, CAG-MIND enables robust zero-shot transfer, data-efficient fine-tuning and cross-center generalization. These findings support domain-aligned multimodal pretraining as a scalable foundation-model paradigm for invasive cardiovascular imaging and future cath-lab decision support.
Modi, D.; Kim, J.; Ye, A.; Eusuff, S.; Ieki, H.; Ambrosy, A. P.; Kwan, A. C.; He, B.; Zou, J.; Cheng, S.; Ashley, E.; Ouyang, D.
Show abstract
Importance: Mobile phone-recorded echocardiogram videos are commonly used in point of care, telemedicine, and resource-limited workflows, but artificial intelligence models for left ventricular ejection fraction (LVEF) estimation have primarily been evaluated on native Digital Imaging and Communications in Medicine (DICOM) videos. Objective: To evaluate whether previously described artificial intelligence models for LVEF estimation retain performance when applied to mobile phone-recorded echocardiographic videos. Design: Multicenter model validation study comparing model-estimated LVEF with clinician reported LVEF. Setting: Three medical centers: Kaiser Permanente Northern California, Beth Israel Deaconess Medical Center through MIMIC-IV-ECHO, and Cedars-Sinai Medical Center. Participants: Source studies with clinician reported LVEF and apical 4-chamber or apical 2-chamber views, yielding 6209 phone-recorded videos from 2648 studies and 2611 patients. Exposures: Mobile phone recording of native echocardiographic videos and fine-tuning of pretrained models using mobile phone-recorded videos from the Kaiser Permanente Northern California training cohort. Main Outcomes and Measures: Mean absolute error in ejection fraction percentage points, R^2 for continuous estimation, and area under the receiver operating characteristic curve for identifying ejection fraction greater than 50%. Results: The study included 6209 mobile phone recorded echocardiographic videos from 2648 studies and 2611 patients; the weighted mean age was 68.4 years, and 1031 patients were male (39.5%). Without phone-video fine-tuning, the primary model achieved a mean absolute error of 7.00 percentage points, coefficient of determination of 0.49, and area under the receiver operating characteristic curve of 0.91 on phone-recorded videos; corresponding native DICOM performance was 6.08 percentage points, 0.60, and 0.93, respectively. On the 2396-video fine-tuning evaluation cohort, fine-tuning improved primary model performance to a mean absolute error of 6.96 percentage points, coefficient of determination of 0.61, and area under the receiver operating characteristic curve of 0.93. Fine-tuning the public EchoNet-Dynamic model improved performance from 9.36 percentage points, 0.37, and 0.84 to 7.86 percentage points, 0.50, and 0.89, respectively. Progressive central zoom preprocessing degraded model performance. Conclusions and Relevance: These findings suggest that artificial intelligence assisted left ventricular ejection fraction estimation from mobile phone-recorded echocardiograms may be feasible when native image export is unavailable, although prospective evaluation is needed before clinical deployment.
Jamthikar, A. D.; Shanmugham, A.; Singh, S.; Radhakrishnan, A.; Dong, J.; Maganti, K.; Yanamala, N.; Sengupta, P.
Show abstract
Background: Left ventricular diastolic dysfunction (LVDD) is a major determinant of heart failure (HF), yet its assessment relies on multiparametric echocardiography, limiting scalability. We previously demonstrated that generative artificial intelligence (AI) can synthesize tissue Doppler imaging (TDI) waveforms from the 12-lead ECG. The growing complexity of candidate architecture creates a need for automated model-discovery frameworks. Objectives: To evaluate agentic AI-based auto-discovery for ECG-based LVDD assessment using either raw ECG or synthetic TDI waveforms. Methods: Two attention-based agentic AI architectures were developed using an automated large language model-driven refinement framework that optimized transfer-learning and multimodal architectures through autonomous proposal, validation, and selection of candidate model configurations. Development was performed in 1,011 paired ECG-echocardiography studies and externally validated in 983 patients using two reference frameworks: (i) data-driven phenogroups and (ii) the 2025 ASE Diastolic Function Guidelines. External validation was performed in CODE-15% (n=219,567) for HF-related mortality and EchoNext (n=35,718) for structural heart disease associations. Results: Despite the modest cohort size, the ECG-based agentic search achieved area under the receiver operating characteristic curve (AUCs) of 0.87 (95% CI: 0.85-0.89) and 0.83 (95% CI: 0.80-0.86) for phenogroup and guideline-based LVDD severity classification. Corresponding AUCs for the synthetic TDI-based model were 0.82 (95% CI: 0.80-0.85) and 0.80 (95% CI: 0.77-0.84), respectively. In large-scale external validation, both models stratified incident HF mortality with subdistribution hazard ratios ranging 5.5 to 9.5 (Gray's p<0.001 for all). Time-dependent discrimination for incident HF mortality exceeded a publicly available convolutional neural network model (ECG2HF) ({Delta}AUC range: +0.14 to +0.20). Both models demonstrated consistent associations with structural heart disease outcomes. Conclusions: Agentic auto-discovery enabled data-efficient assessment of LVDD from surface ECG by combining physiologically informed transfer learning with autonomous architecture optimization, achieving robust external generalizability. This approach may facilitate broader access to diastolic function assessment beyond conventional echocardiography.
Kaplan, T.; Ramirez, J.; Young, W. J.; Sanghvi, M. M.; Madrid, J.; Naderi, H.; Shah, R.; Minchole, A.; Orini, M.; Tinker, A.; Lambiase, P. D.; Munroe, P. B.; van Duijvenboden, S.
Show abstract
Electrocardiographic (ECG) interval measurements underpin clinical decision-making and large-scale cardiovascular research, yet existing automated methods are often developed using small, heterogeneous datasets with limited expert annotation and uncertain generalisability to population cohorts. We developed a deep learning framework for automated PR, QRS, and QT interval estimation and established a large expert-curated reference dataset using UK Biobank (UKB) ECGs. The reference dataset comprises 11,330 lead-level annotations from 12-lead ECGs in 1,030 randomly selected UKB participants, generated using a standardised annotation protocol with independent expert review. A 1D convolutional neural network was trained to segment ECG waveforms and derive PR, QRS, and QT intervals. Performance was evaluated against expert annotations, UKB CardioSoft measurements, an open-source signal-processing toolbox, and a wavelet-based delineation method. Clinical validity was evaluated through associations with incident atrial fibrillation and major adverse cardiovascular events (MACE). Inter-observer agreement was high (ICC 0.81 - 0.97). In a held-out test set, the deep learning model achieved mean absolute errors of 7.7 ms (PR), 7.5 ms (QRS), and 4.9 ms (QT), outperforming all comparator methods, with minimal bias relative to expert annotations. Review of distributional outliers confirmed >80% validity for most interval measurements. Among 46,749 participants with follow-up (median 4 years), prolonged QTc derived by the deep learning model showed stronger associations with incident MACE (hazard ratio 2.9, 95% CI 2.1 - 4.0) than wavelet-based measurements (1.7, 1.4 - 2.0) or CardioSoft (1.1, 0.9 - 1.4). This expert-curated reference dataset and validated deep learning framework provide a scalable foundation for reproducible ECG phenotyping in UKB and beyond.
Sahashi, Y.; Choi, D.; Ieki, H.; Vukadinovic, M.; Rawlani, M.; Kwan, A. C.; Cheng, S.; Ouyang, D.
Show abstract
Summary Background A comprehensive transthoracic echocardiogram involves the assessment of over 70 parameters, placing a substantial burden on sonographers and physicians for manual annotation with considerable inter-observer variability. Prior open-source segmentation models have largely addressed 2D B-mode ventricular function, leaving a gap in the spectral Doppler and atrial measurements required for valvular and diastolic assessment such as velocity-time integral (VTI) and atrial chamber size. Methods In this retrospective multi-cohort study, we developed EchoNet-Segmentation, comprehensive task-specific deep learning segmentation models for left and right atrial area and VTI Doppler measurements. Training used 186,712 sonographer-annotated images from 93,978 studies (56,855 patients) at Cedars-Sinai Medical Center (CSMC). Performance was evaluated on a held-out CSMC test set, a CSMC temporal split, an external Kaiser Permanente Northern California cohort, and the public MIMIC-Echo dataset. Findings On the CSMC held-out test set, our AI models showed strong agreement with sonographer measurements, with R2 of 0.817-0.882 and mean absolute error (MAE) of 1.13-3.80 cm for automated VTI measurements, and R2 of 0.675-0.747 and MAE of 2.48-2.52 cm2 for left and right atrial area segmentation. Performance was consistently confirmed on the CSMC temporal split (VTI: R2 0.606-0.866, atrial area: R2 0.694-0.705) and on the KPNC external cohort (VTI: R2 0.575-0.859, atrial area: R2 0.803-0.876), on the MIMIC-Echo dataset. Robustness was demonstrated on a different vendor's machines and across subgroups. EchoNet-Segmentation outperformed an open-source medical image foundation model with bounding-box, point prompt configurations on R2, MAE, and Dice score on both held-out test dataset and MIMIC apical four-chamber data. Interpretation EchoNet-Segmentation is the first open-source framework that delivers accurate, generalizable automated measurement across several key routine echocardiographic parameters, supporting end-to-end automation of clinically important echocardiographic assessments. Public release of model weights, code, and demonstration tools can facilitate reproducibility, research use and clinical deployment.
Shimada, T.; Kodera, S.; Sawano, S.; Guan, J.; Saitoh, W.; Wakasa, S.; Ito, S.; Yanagishita, T.; Hayashi, Y.; Shibata, A.; Ito, A.; Otsuka, K.; Higashikuni, Y.; Okamura, H.; Tsujita, K.; Node, K.; Yamaguchi, O.; Makimoto, H.; Kabutoya, T.; Imai, Y.; Nakayama, M.; Sato, H.; Fujita, H.; Kohro, T.; Matoba, T.; Takeda, N.; Fukuda, D.; Nagai, R.
Show abstract
Background: Aortic stenosis (AS) is a progressive valvular disease associated with poor prognosis once symptoms develop, yet routine echocardiographic screening is impractical. While artificial intelligence (AI)-based electrocardiogram (ECG) models have shown promise for AS detection, it remains unclear whether they primarily reflect conventional left ventricular hypertrophy (LVH) voltage criteria or capture additional ECG features. Methods and Results: We developed a deep learning model using 244,816 ECGs from 51,713 patients across six academic institutions in Japan (CLIDAS database). AS labels were derived from inpatient Diagnosis Procedure Combination (DPC) codes. The model achieved an area under the receiver operating characteristic curve (AUC) of 0.849 (95% confidence interval 0.832-0.865) in the independent test cohort, with consistent performance across institutions, sex, and age. At a threshold of 0.1, sensitivity was 79.1%, specificity was 73.9%, and negative predictive value (NPV) was 98.0%. Conventional LVH voltage criteria (Sokolow-Lyon AUC 0.706; Cornell AUC 0.692) showed lower performance, and adding them to the AI model conferred no incremental benefit (AUC 0.849 vs. 0.847). Gradient-weighted class activation mapping (Grad-CAM) revealed predominant attention around QRS complexes in limb leads, beyond regions typically assessed in LVH evaluation. Conclusions: This multicenter AI-ECG model demonstrated strong discrimination for AS and captured ECG features beyond conventional LVH voltage criteria. The high NPV supports its use as a rule-out pre-screening tool.
Lee, H. S.; Kang, S.; Lee, M. S.; Pandey, A.; Kim, M.; Jang, J.-H.; Jo, Y.-Y.; Lim, J.; Son, J. M.; Kim, K. S.; Kwon, J.-m.; Lee, S.-P.; Kim, K.-H.
Show abstract
Background Structural heart disease (SHD) drives heart failure and cardiovascular mortality but remains underdiagnosed, and echocardiography is limited as a population-level screening tool. Objectives We evaluated whether a composite artificial intelligence-enabled electrocardiogram (AI-ECG), combining independently developed models for left ventricular systolic (LVSD) and diastolic dysfunction (LVDD), identifies prevalent and predicts incident SHD across diverse populations. Methods In this multinational cohort study, detection was assessed cross-sectionally in a Korean clinical cohort (Incheon Sejong Hospital) and a US dataset (Columbia University Irving Medical Center), and incident risk was assessed in the Korean cohort and the UK Biobank among individuals without baseline SHD or heart failure. Adults with paired ECG and echocardiography were analyzed for detection, with the composite defined as positive on either model. SHD comprised reduced left ventricular ejection fraction, moderate or severe valvular disease, left ventricular hypertrophy, or pulmonary hypertension. Detection was assessed by sensitivity and specificity, and incident risk by Cox models and the C statistic. Results Among 46,082 and 36,286 participants in the two detection cohorts, the composite detected SHD with sensitivity of 71.8% and 76.1% and specificity of 88.3% and 70.1%, with positivity across all phenotypes. Among at-risk individuals, composite positivity was associated with incident SHD (hazard ratios, 3.75 and 2.75), with C statistics of 0.69 to 0.78. Conclusions A composite AI-ECG identified prevalent and predicted incident SHD across multinational cohorts, capturing signals beyond its training targets and supporting its potential as a scalable cardiovascular screening tool; whether ECG-based risk stratification improves outcomes requires prospective evaluation.
Tiruwa, K. R.; Ghimire, A.
Show abstract
Consumer wearable devices increasingly use single-lead electrocardiograms (ECGs) for cardiac monitoring, but these signals contain substantially less spatial information than the clinical 12-lead standard. Whether this reduction dispro- portionately affects older adults, who often present with more complex cardiac conditions, remains poorly understood. In this study, we evaluated the impact of lead reduction on AI-ECG diagnostic performance across age groups. A 1D resid- ual neural network was trained on 21,091 PTB-XL ECG recordings spanning five diagnostic superclasses and assessed using 12-, 6-, 2-, and 1-lead configurations. Under the full 12-lead setting, model accuracy declined from 84.5% in patients younger than 40 years to 66.2% in patients aged 75 years or older. Progressive lead reduction further widened this gap. Under the 1-lead configuration, accuracy decreased by 14.1 percentage points in the 75+ group but by only 0.4 percent- age points in the <40 group, representing an approximately 40-fold differential degradation confirmed by three independent statistical tests (all p < 0.0001). Older adults also exhibited greater multi-condition diagnostic complexity, pro- viding a plausible explanation for their increased vulnerability to information loss. External validation on the MIT-BIH Arrhythmia Database confirmed cross- dataset model stability. These findings suggest that age-stratified performance reporting should be a minimum standard in wearable AI-ECG validation and regulatory assessment.
Alavi, R.; Li, J.; Matthews, R. V.; Pahlevan, N. M.; Kloner, R. A.; Gharib, M.
Show abstract
The electrocardiogram (ECG) contains rich nonlinear and non-stationary dynamic information that is only partly captured by conventional ECG interpretation and beat-to-beat metrics, and is increasingly analyzed using black-box artificial intelligence models that often lack interpretability. Here, we introduce the ECG time-frequency "eyeball", an interpretable framework that transforms a brief single-lead ECG recording into a geometric signature and a set of low-dimensional rotational and geometrical features using empirical mode decomposition and Hilbert-based analytic signal mapping. In 30-second lead I-equivalent recordings from 170 healthy subjects and 80 patients with acute myocardial infarction (AMI), the proposed "eyeball" metrics significantly differentiated groups, with AMI associated with higher rotational frequency metrics, lower envelope metrics, and displaced centroid location. Representative examples revealed a coherent morphologic spectrum from normal patterns to geometries consistent with myocardial ischemia, injury, and infarction. The representation remained stable across recording windows from 30 seconds to 5 minutes, and individual "eyeball" features achieved areas under the receiver operating characteristic curve (AUCs) of up to 0.78 for AMI detection. These findings suggest that the ECG time-frequency "eyeball" condenses clinically relevant nonlinear ECG dynamics into an interpretable representation that may reveal hidden AMI signatures, complement conventional ECG interpretation, and provide a foundation for accessible single-lead cardiovascular screening using future smart wearables.
Garcia, N. M.
Show abstract
Conventional electrocardiography is highly effective for waveform and rhythm diagnosis, but it is less suited to showing how the internal shape of hundreds or thousands of consecutive heartbeats changes over time. We introduce FOXTAIL, a complementary view that represents each cardiac cycle as an ordered sequence of changes in signal direction. Overlaying these sequences in a fixed visual field makes beat-to-beat organization visible and allows the density, size, stability, and scale persistence of those changes to be measured. We evaluated the representation in recordings containing normal sinus rhythm, paroxysmal atrial fibrillation, severe heart failure, ventricular tachyarrhythmia, and controlled electrode-motion noise. Paired recordings showed that FOXTAIL descriptors can reveal within-person state changes that are not conveyed by a single average beat. The noise and pre-fibrillation analyses also showed that a dense event pattern is not automatically equivalent to physiological complexity, measurement artifact, or impending disease. FOXTAIL is therefore not proposed as a replacement for the diagnostic ECG or as a new classifier, but as an observation and measurement domain for asking a more basic question: how is the electrical organization of the heart changing from one beat to the next, and which of those changes persist across scale?
Lampadarios, T.; Karathanasis, N.; Antartis, R.; Pfeifer, B.; von Lewinski, D.; Sourij, H.; Spyrou, G. M.; Oulas, A.
Show abstract
Acute myocardial infarction (MI) is a major precursor to heart failure (HF), yet few biomarkers are routinely used to predict post-MI HF, and limited therapeutic options exist to prevent its development. Furthermore, identifying patients at extremely high risk of recurrent MI remains challenging. These gaps highlight the need for improved biomarkers, therapeutic targets, and computational approaches for risk assessment and treatment-response prediction. To address risk assessment, we developed a systems bioinformatics (SB), graph-based framework representing patient information as personalized networks and integrating omics, clinical, and molecular prior-knowledge data. Graph neural network (GNN) machine learning (ML) models were compared with conventional ML approaches. Two large-scale public plasma proteomic datasets were used to predict post-MI HF. To investigate treatment response, regression models were applied to longitudinal clinical data from >400 hospitalized patients enrolled in the EMMY trial evaluating empagliflozin. ML-driven feature selection identified proteins and clinical parameters with the greatest predictive value. The graph-based framework demonstrated strong and consistent performance across independent post-MI cohorts. GNN models outperformed conventional approaches, including generalized linear models and XGBoost, particularly when attention mechanisms were incorporated. Using biomarker panels alone, the best GNN achieved an external test AUC of 0.82, compared with 0.77 for the best conventional ML model. When biomarkers were combined with clinical and demographic variables, GNN and conventional ML models achieved AUCs of 0.80 and 0.77, respectively. Regression models also showed promise for predicting biomarker changes associated with treatment response, with the best model achieving a test RMSE of 0.56. Feature-importance analysis identified NT-proBNP (NPPB), cardiac troponins (TNNI3/TNNT2), and prior HF history as the most influential predictors, consistent with established clinical evidence. Overall, these findings support graph-based ML and regression analysis as promising approaches for improving post-MI HF risk prediction and therapeutic response and identifying clinically relevant markers.